Papers with data generation framework

8 papers
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)

Copied to clipboard

Challenge: Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems.
Approach: They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples.
Outcome: The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache .
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs’ Memorization (2025.acl-long)

Copied to clipboard

Challenge: UnSeenTimeQA is a data contamination-free time-sensitive question-answering benchmark.
Approach: They propose a data contamination-free time-sensitive question-answering benchmark that avoids web-searchable queries grounded in the real world.
Outcome: The proposed benchmark avoids web-searchable queries grounded in the real world and enables on-demand generation of new samples, mitigating the risk of data leakage.
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking (2024.eacl-long)

Copied to clipboard

Challenge: In-context learning with Large Language Models (LLMs) is a promising avenue of research in Dialog State Tracking (DST).
Approach: They propose a data generation framework tailored for Dialog State Tracking that uses large language models to synthesize natural, coherent, and free-flowing dialogues with DST annotations.
Outcome: The proposed framework improves joint goal accuracy by 4-5% over the zero-shot baseline on MultiWOZ 2.1 and 2.4.
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models often exhibit gender bias, resulting in unequal treatment of male and female subjects across contexts.
Approach: They propose a framework that encourages exploratory thinking in large language models . the framework generates story pairs featuring male and female protagonists in structurally identical scenarios .
Outcome: The proposed framework reduces gender bias while preserving or even enhancing general model capabilities.
Towards Better Hierarchical Text Classification with Data Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve hierarchical text classification are expensive and lack high-quality labeled data.
Approach: They propose a hierarchical text classification framework that can achieve both label controllability and text diversity by extracting high-quality hierarchic label information.
Outcome: The proposed method can achieve label controllability and text diversity by extracting high-quality hierarchical label information.
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to manage topic shifts within on-topic dialogues are limited in their ability to generate training datasets.
Approach: They propose a data generation framework that automatically generates conversational question-answering datasets with natural topic transitions by leveraging relationships between entities in a knowledge graph.
Outcome: The proposed framework generates conversational question-answering datasets with natural topic transitions and proves its effectiveness in generating dialogues with topic shifts.
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following (2025.acl-long)

Copied to clipboard

Challenge: Multi-document (MD) processing is crucial for LLMs to handle real-world tasks such as summarization and question-answering across large sets of documents.
Approach: They propose a framework that generates high-quality synthetic MD instruction data over sets of articles via targeted prompts.
Outcome: MDCure generates high-quality synthetic MD instruction data over sets of articles . evaluations show it improves over pre-trained models by up to 75.1% .
PhaseMI: A Motivational Interviewing Dataset for Enhancing Phase Progression in LLM-based Counseling (2026.findings-acl)

Copied to clipboard

Challenge: Existing MI datasets do not explicitly model structured progression of MI phases, which is essential for effective and goal-oriented counseling.
Approach: They propose a phase-structured MI dataset with a data generation framework that employs therapist, client, and supervisor LLMs to explicitly control phase transitions.
Outcome: The proposed model achieves 12.3% better coverage of MI phases, 37.6% in guiding, and 61.1% in choosing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations